A Parallel Learning Algorithm for Text Categorization on PIRUN Beowulf Cluster
نویسندگان
چکیده
Text categorization is the process of classifying documents into predefined categories or classes based on their content. Since text data rapidly increase on the Internet, the scalability of the algorithm is required to handle such massive data. In this paper, we propose a parallel learning algorithm for text categorization based on the combination of the Expectation-Maximization (EM) algorithm and the naive Bayes. Our experiment performed on a 72 nodes Beowulf cluster called PIRUN. The preliminary experimental results show that our parallel implementation has reasonable speedup characteristics.
منابع مشابه
Improving the Operation of Text Categorization Systems with Selecting Proper Features Based on PSO-LA
With the explosive growth in amount of information, it is highly required to utilize tools and methods in order to search, filter and manage resources. One of the major problems in text classification relates to the high dimensional feature spaces. Therefore, the main goal of text classification is to reduce the dimensionality of features space. There are many feature selection methods. However...
متن کاملBuilding a Large Scalable Internet Superserver for Academic Services with Linux Cluster Technology
With the speed and bandwidth offered by the next generation Internet technology, there is a need for large and scalable Internet server that can provides an adequate computing power and storage for the new generation Internet applications. This requires a huge investment in a very large and expensive commercial server system. Recently, the emergence of Linux PC clustering or so-called Beowulf C...
متن کاملParallel Nearest Neighbour Algorithms for Text Categorization
In this paper we describe the parallelization of two nearest neighbour classification algorithms. Nearest neighbour methods are well-known machine learning techniques. They have been successfully applied to Text Categorization task. Based on standard parallel techniques we propose two versions of each algorithm on message passing architectures. We also include experimental results on a cluster ...
متن کاملParallelization of Noise Reduction Algorithm for Seismic Data on a Beowulf Cluster
This paper presents the parallelization of a sequential noise reduction algorithm for seismic data processing into a parallel algorithm. The parallel algorithm was developed using C language with the utilization of the Message Passing Interface (MPI) library. The proposed algorithm has been implemented on an experimental Beowulf cluster which consists of 12 nodes operating on Linux Ubuntu platf...
متن کاملSolving Traveling Salesman Problem on Cluster Compute Nodes
In this paper, we present a parallel implementation of a solution for the Traveling Salesman Problem (TSP). TSP is the problem of finding the shortest path from point A to point B, given a set of points and passing through each point exactly once. Initially a sequential algorithm is fabricated from scratch and written in C language. The sequential algorithm is then converted into a parallel alg...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2005